Wednesday, April 19, 2006

Classification of XML data

As part of the 7DS community extension, I will have to classify shared community XML objects using some sort of XML schema, RDF or RDF Schemas. We looked at OWL today, but that looks fairly complicated, so we may just go ahead and use RDF and RDF schemas.

Friday, April 14, 2006

PHP and web (http) servers on Windows Mobile / PocketPC

As the scope of the 7DS project expands, I'm coming up against building more involved 7DS web applications. The search and multicast engines were C programs that produced binary CGI executables, but building more involved community-based web systems is going to be hard. :(

In looking to build the new applications in PHP, I was asked to see if PHP might be supported on Windows Mobile (one of our near-future development platforms) and I found it interesting that Windows Mobile SDK has its own HTTP/web server.

Even more interesting, there have been problems porting the Zend PHP engine to Windows Mobile, so there is an alpha version of a ground-up PHP engine built for Windows Mobile with its own web server.

Friday, March 31, 2006

Always use sizeof() when malloc()ing

I've had some real big problems implementing my webpage retreiver, and I fixed it once I realized that I had to include the sizeof() everytime I malloc()d.

So remember to use this everytime you malloc, say a string:


newstring = malloc ((strlen(oldstring)+1) * sizeof(char));


Remember to add 1 as well, just as I did: C strings terminate with a '\0', which is an extra character.

Thursday, March 30, 2006

Memory leak for null string assignment

This is really wierd... the following piece of code in my 7DS system results in a huge memory leak, gobbling up memory really fast.


if (0 == hits) {
// No results, empty xmlResults and return
sprintf (xmlResult, "");
return 0;
}


Disabling it solves the memory problem - I wonder why?

Tuesday, March 28, 2006

malloc() and free() error for dynamic strings solved

OK, I am a newbie at dynamic strings in C, so please forgive my silliness.

I have been getting "*** glibc detected *** free(): invalid next size (fast)" errors in my application that has to create dynamic path names and couldn't figure it out.

Finally I did a Google search and found the solution here: http://www.eskimo.com/~scs/cclass/int/sx7.html

Guess what I had done? malloc()d the string to use one less character than needed like this:
escapedURL = malloc (strlen (URL));

This is CORRECT:
escapedURL = malloc (strlen (URL) + 1);

Because C strings end with a '\0' character.

Monday, March 27, 2006

Parsing filename and directories out of given path in C

Sample code for how to parse directory structure and path names for a given path string in C. I will use this for the webpage retreiver project I am working on.


#include <string.h>
#include <stddef.h>
#include <libgen.h>
#include <stdio.h>

/* This program parses the path given in the argument into directories
* and filename, creates the directory structure and creates an
* empty file as well */
int
main (int argc, char **argv)
{
const char delimiters[] = "/\\"; /* File delimiters */
char *token, *oldtoken, *cp;
const char program_directory = getcwd (NULL, 0); /* Program directory */

/* Create copy of path string */
cp = malloc (strlen (argv[1]));
strcpy (cp, argv[1]);

/* Split string into tokens */
token = strtok (cp, delimiters);

printf ("Directory = ");

/* While token is not NULL */
while (1)
{
oldtoken = malloc (strlen (token));
strcpy (oldtoken, token); /* Copy current token */
token = strtok (NULL, delimiters); /* Go to the next token */
/* If nexxt token is NULL, it is the last part and assumed to
* be a filename */
if (token == NULL)
{
printf ("\nFilename = %s\n", oldtoken);
/* Create an empty file of that name */
FILE *fp;
fp = fopen (oldtoken, "w");
fclose (fp);
break;
}
/* Otherwise it is a directory */
else
{
printf ("%s ", oldtoken);
mkdir (oldtoken, 0755); /* Create the directory */
chdir (oldtoken); /* Go there */
}
free (oldtoken);
}
printf ("\n");
/* Free any memory elements */
free (cp);
oldtoken = NULL;

chdir (program_directory);
}

Friday, March 17, 2006

URL or URI parsing in libxml and c

I found out that libxml has URI functions that will allow you to parse URIs using libxml.

Here's an example program:


#include <stdio.h>
#include <libxml.h>

int main(int argc, char **argv) {

/* Create a null URI */
xmlURIPtr url = xmlCreateURI ();

/* Parse the user input URI */
url = xmlParseURI ( (argc <= 1) ? "http://www.theepochtimes.com/" : argv[1]);

/* Print all the respective information */
printf ("scheme = %s\n", url->scheme);
printf ("opaque = %s\n", url->opaque);
printf ("authority = %s\n", url->authority);
printf ("server = %s\n", url->server);
printf ("user = %s\n", url->user);
printf ("port = %d\n", url->port);
printf ("path = %s\n", url->path);
printf ("query = %s\n", url->query);
printf ("fragment = %s\n", url->fragment);
printf ("cleanup = %d\n", url->cleanup);
}

Parsing HTML using tidy and tidylib

It's so hard to find a C program on the web that can parse HTML! Yes, you can find parsers written in Perl and other languages, but not C!

So I might as well share what I've learnt so far. I am making the 7DS HTML parser in libxml, but I experimented using tidy and tidylib as well, and here's how the code for that looks:


#include <tidy.h&rt;
#include <buffio.h&rt;
#include <stdio.h&rt;
#include <errno.h&rt;

/**
* Dump the list of nodes and their attributes
* Modified from tidylib documentation
*/
void dumpNode( TidyNode tnod, int indent )
{
TidyNode child;

for ( child = tidyGetChild(tnod); child; child = tidyGetNext(child) )
{
ctmbstr name = tidyNodeGetName( child );
if ( !name )
{
switch ( tidyNodeGetType(child) )
{
case TidyNode_Root: name = "Root"; break;
case TidyNode_DocType: name = "DOCTYPE"; break;
case TidyNode_Comment: name = "Comment"; break;
case TidyNode_ProcIns: name = "Processing Instruction"; break;
case TidyNode_Text: name = "Text"; break;
case TidyNode_CDATA: name = "CDATA"; break;
case TidyNode_Section: name = "XML Section"; break;
case TidyNode_Asp: name = "ASP"; break;
case TidyNode_Jste: name = "JSTE"; break;
case TidyNode_Php: name = "PHP"; break;
case TidyNode_XmlDecl: name = "XML Declaration"; break;

case TidyNode_Start:
case TidyNode_End:
case TidyNode_StartEnd:
default:
assert( name != NULL ); // Shouldn't get here
break;
}
}
assert( name != NULL );
char whitespace[indent];
memset (whitespace, ' ', indent);
whitespace[indent-1] = '\0';
// printf( "%sNode: %s\n", whitespace, name );

/* Get the first attribute for all nodes */
TidyAttr tattr = tidyAttrFirst (child);
while (tattr != NULL) {
/* Print the node and its attribute */
printf ("%s %s %s= %s\n", whitespace, tidyNodeGetName (child), tidyAttrName (tattr), tidyAttrValue (tattr));
/* Get the next attribute */
tattr = tidyAttrNext (tattr);
}
dumpNode( child, indent + 4 );
}
}

/* Dump the whole document */
void dumpDoc( TidyDoc tdoc )
{
dumpNode( tidyGetRoot(tdoc), 0 );
}

/* Dump only the body */
void dumpBody( TidyDoc tdoc )
{
dumpNode( tidyGetBody(tdoc), 0 );
}

int main(int argc, char **argv )
{
/* Input file: Either the first argument or "../test.html" */
const char* input = (argc > 1) ? argv[1] : "../test.html";
TidyBuffer output = {0};
TidyBuffer errbuf = {0};
int rc = -1;
Bool ok;

TidyDoc tdoc = tidyCreate(); // Initialize "document"
printf( "Tidying:\t%s\n", input );

ok = tidyOptSetBool( tdoc, TidyXhtmlOut, yes ); // Convert to XHTML
if ( ok )
rc = tidySetErrorBuffer( tdoc, &errbuf ); // Capture diagnostics
if ( rc >= 0 )
/* Read from the HTML file */
rc = tidyParseFile( tdoc, input ); // Parse the input
if ( rc >= 0 )
rc = tidyCleanAndRepair( tdoc ); // Tidy it up!
if ( rc >= 0 )
rc = tidyRunDiagnostics( tdoc ); // Kvetch
if ( rc > 1 ) // If error, force output.
rc = ( tidyOptSetBool(tdoc, TidyForceOutput, yes) ? rc : -1 );
if ( rc >= 0 )
rc = tidySaveBuffer( tdoc, &output ); // Pretty Print

if ( rc >= 0 )
{
if ( rc > 0 )
printf( "\nDiagnostics:\n\n%s", errbuf.bp );
printf( "\nAnd here is the result:\n\n%s", output.bp );
}
else
printf( "A severe error (%d) occurred.\\n", rc );

tidyBufFree( &output );
tidyBufFree( &errbuf );

/* Now parse and print the tags in the HTML document */
dumpDoc (tdoc);

tidyRelease( tdoc );
return rc;
}

Tuesday, March 14, 2006

Webcrawler using libxml, libcurl and tidy

Contrary to my writeup in the last post about how wget might be the best way to webcrawl and fetch files to a local cache, my thoughts now are different.

You can use the following libraries to build a decent webcrawler:

1. Tidy: Use tidylib to clean up your HTML pages and make them XHTML. tidylib's webpage has sample code that is good enough for converting HTML to XHTML - just make sure you save to a file using tidySaveFile().

libxml has problems parsing HTML, even if used with xmlRecoverFile() rather than xmlParseFile().

2. libxml: Parse the XHTML, get all elements' attributes (and any other URLs you need) and pass on the URLs to libcurl to download. Need I say more?

Well, actually I should. libxml is a little hard to understand from the API, and sample code to do what you want is hard to find. I had to do quite a bit of searching, looking up sample programs, and then reading the API to figure out how things worked.

3. curl: Or rather libcurl. To retrieve files from the Net. Again, need I say more?

Life would have been simpler if curl had a recursive download function ... or wget had a library I could use ... but then, that's why we computer engineers and students have a life!

Tuesday, February 28, 2006

wget to create local cache of webpage

Here's how to use wget to create a local cache of a webpage:
wget -r -l 1 –p –-convert-links http://www.yourdomain.com

For more information: http://www.devarticles.com/c/a/Web-Services/Website-Mirroring-With-wget/1/

Tuesday, February 14, 2006

Very nice and detailed documentation on creating RPMs:
http://www.gurulabs.com/GURULABS-RPM-LAB/GURULABS-RPM-GUIDE-v1.0.PDF

It includes information not present in the official RPM FAQ at http://www.rpm.org/

Monday, February 06, 2006

Processes in C: /dev/urandom C code, forking several child processes...

Hi,

Sorry, I've been busy with courses AND research lately, it 's given me less time to post on the blog.

For an OS course, we have to do the Miller-Rabin test for primality, and it involves forking atleast 3 child processes; getting random numbers from /dev/urandom, etc.

Here is some code for having 2 or more child processes:
http://www.csl.mtu.edu/cs4411.ck/www/NOTES/process/fork/fork-03.c

For reading from /dev/urandom (and why that's a good idea):
http://www.kalyanvarma.net/tech/security/random.html

Tuesday, December 13, 2005

Got Darwin Streaming Server running in Linux

Buggered if I know why ... but I installed Darwin Streaming Server without a root account and changed the settings and other files as needed.

Then, yesterday, it worked when I started it, but then , I shut it down and couldn't start it again.

Then, after many tries, it finally worked when I chdir'd to /homes/streaming/usr/local/sbin and ran this command:
./DarwinStreamingServer -c /homes/streaming/etc/streaming/streamingserver.xml

But it wouldn't run from any other directory!

The web admin page still does not work; wonder why.

Monday, December 12, 2005

PhD qualifiers & Installing Darwin without root account

Sorry, I've disappeared for a while since I'm preparing for the EE department's PhD qualifying exams.

I still have to install Darwin Streaming Server for the CS department's videos. My colleague Salman tipped me off to this great link on how to install Darwin without a root or su account.

But the new Darwin SS has some changes from the link that is posted. You might want to do the following instead:
- Edit the Install script so that the directories are your directories
- In Install script, comment the lines to create user qtss and the chmod lines
- Change the streamingserver.xml he mentions to update the directory links
- The new streamingadminserver.pl has default configuration values in the script itself, so you have to change those values rather than change an external config file.

If you are using vi, you can do the replacements by issuing the following commands:

:%s/\/usr\//\/homes\/streaming\/usr\//g
:%s/\/var\//\/homes\/streaming\/var\//g
:%s/\/etc\//\/homes\/streaming\/etc\//g

If you are editing streamingserver.pl, you ONLY need to do that for lines 230 - 270 in the current version, because only those are Linux-specific config settings. (Replace % with 230, 270)

I basically diff'd the config and script files he had with mine and decided to just play it safe and modify the default files.

Other than that, you should hear more from me after I'm done with my PhD quals in early January ... but at that point, I'll be taking Operating Systems and a seminar course for the Spring 2006 semester.

Monday, November 07, 2005

NetSurf - Source code for browser

Right after I posted the last post about lack of open-source documentation for HTML parsing, I stumbled across this neat browser for RISC OS: NetSurf. And the full source code is neatly documented!

I plan to go over the code for NetSurf in building my new app.

libxml for parsing HTML

Long time no see. Have been busy with preparing for my EE PhD qualifiers!

Well anyway, now I am to create a HTTP-retreiver-&-parser to get files for the 7DS multicast query system. It now has to not only get the result-set for the 7DS queries, but should also get the files themselves, as well as associated elements, such as images, etc.

I found several HTML parsers for C (after long searches) such as ekhtml (nil documentation), tidy (library does not build properly) and LibWWW (supposed to be very complicated) ... and have settled on using LibXML's inbuilt HTML parsing tools.

Sad that open source code has very little documentation ... hey, but neither does 7DS yet!

Monday, October 10, 2005

Reformatting source

I came back from India 2 weeks back, and am spending time reformatting my C source code.

I am going to use a tool called DOC++ for documentation. Also, I am going to split my code in such a way that commonly used functions go into a "library" file.

Thursday, August 11, 2005

Wireless card woes

I'm still having problems getting the Senao wireless card to work on WRAP. :(

I am coming to the conclusion that maybe ... there is something wrong with the device itself?

I finally managed to get a copy of lspci for Bering and ran it on WRAP - I got it here http://fritzfam.com/brad/leaftmp/

But now that I am running lspci, I don't see any wireless card on the PCI list!! :(

My last hope - that it is a PCI card with a PCMCIA - PCI bridge ... but there's not very much hope.

7DS works on WRAP

I got 7DS working on WRAP now!

I've mostly focussed on getting the wireless card to work with WRAP for the past few weeks... and in the meantime, I did some work on getting 7DS running on the WRAP board.

Here are what I did:
- Put all the 7DS binaries into 7ds.lrp
- Ran the query_receiver and it received multicast packets sent by another machine
(This was a little wierd; multicast did not work one day, but it worked the next.)

- Copied the directory structure using /home/local/sumans/ (same as on the development machines) and put this into 7ds_cach.lrp
- Copied the db.7ds.queries sqlite3 database to the appropriate place on the WRAP board

and it worked!

Also, I got the 7DS web interface and CGIs working as well. Here's how that was done:
- The website and CGIs are already in 7ds_cach.lrp
- I had to modify the mhttpd setting on WRAP and add this line:
dir = /home/sumans/local/www

and that worked as well!

Monday, July 25, 2005

natsemi network card on WRAP

Well, I got the NatSemi network card working on the WRAP board after a lot of confusing tries!

First of all, I had to uncomment the lines for "crc32" and "natsemi" to load the modules at start up - which I forgot to do.

Once I did that, I got some "unresolved symbol" error messages. A post to the LEAF mailing list revealed that I might have different versions of the kernel and modules - which was the case! I had downloaded the latest version of kernel and modules, but used an older version of the WRAP-specific kernel ... which is where the problem was.

I then updated the WRAP board to use the same versions of the kernel and modules, and it worked!

Not only did the natsemi driver work, but I also got the 7DS multicast network working today! It didn't work last week (CRF told me the systems I was using were behind "different switches") but they sure do work now.