Showing posts with label ek$iAPI. Show all posts
Showing posts with label ek$iAPI. Show all posts

Friday, August 31, 2007

boşları alalım

asansörlerden nefret ediyorum. hadi biraz yumuşatalım, hazzetmiyorum. diyeceksiniz ki neden? çünkü zaman ve mekan bakımından fena halde verimsizler. sabah, öğle, akşam hiç farketmiyor; herhangi bir kata gitmek için genellikle 1 dakikadan fazla beklemek zorunda kalıyorum, kalıyoruz. kule'deki asansörlerin işletmesinde kullanılan sistem nasıl bir algoritma kullanıyor bilmem, gelişkin birşey olması gerektiği de muhakkak, ama ben beklerken koridorda tap dancing harikaları yaratacaksam neye yarar? bu işe biraz kafa yormak lazım. yoralım:

  • eğer asansör sayımız birden fazlaysa bunları sektörlere ayırabiliriz. her kabin öncelikle belli kat aralıklarına hizmet verir, ama kendi sektörünün dışına çıkabilir. boşta kaldığı zaman da kendi sektörünün dışındaysa sektörü içindeki en yakın kata hareket eder.
  • kabin içi butonlarına basma istatistiği tutabilen bir sistem kurulmalı. atıyorum, 9. kattan en çok 12. kata mı çıkılıyor? o zaman buraya 9. ve 12. kata en yakın asansörü gönderebiliriz. bu istatistikler sektör sınırlarının belirlenmesinde de yardımcı olacaktır.
  • katlarda asansör çağırma talepleri zamana göre sıralanmalı, önce çağırana asansör daha erken gitmeli. tabi yol üstünde daha sonra çağıran varsa onları es geçmek de olmaz. yine de hiç yoktan kaynak kıtlığı yaratmamak lazım.
  • katlarda bekleyen kişi sayısı, daha doğrusu bekleyenlerin kilo cinsinden toplam ağırlığı da önemli. toplam büyükse (ve eğer grup fazla beklememişse) daha boş bir kabin gelene kadar grup bekletilebilir. daha da abartalım ve ağırlık değişimini izleyelim. böylece bizim grup olarak addettiğimiz "kütle"nin aslında daha ufak gruplardan oluşup oluşmadığını anlayabiliriz. çağırma taleplerini de bu grafikle beraber inceleyip ortaya karışık birşeyler yapılabilir. ufaktan bir sırt çantası problemi uygulamasına dönebilecek bir durum.
  • kabinin bir yöne giderkenki hızı da denklemimize dahil etmemiz gereken bir değişken. eğer kabin benim bulunduğum yöne gelirken ama benim için duramayacak durumdayken asansörü çağırırsam talebim hemen değerlendirilmeyecek, kuyruğa atılacaktır.
yordum. bunları harmanlayıp bi'şeyler çıkartmak fena olmazdı.

sonracığıma, bitirme projemi .net framework'teki yeniliklere uydurmak (öncekini framework 1.1 için yazmıştım, generics falan hak getire) ve harala gürele yazılmış kodu temizlemek için tekrar yazıyorum. ek$i'den entry çekmek için kullandığım IE tabanlı web scraper'ımın kullandığı ekşiAPI'dan başladım. CodeProject'te şu elemanın kullandığı yöntem gayet hoşuma gitti, o yüzden IE otomasyonun atıp buna benzer bir yapıya geçeceğim. ayrıca, sadece ek$i odaklı bir ürün olmayacak. mesela bir class ve okunacak html sayfasındaki bilgilerin verilen class'ın hangi alanlarına dolacağını belirten bir xml dosyasını kullanarak veriyi daha programlanabilir bir şekilde sunacak. değindiğim şeyler yeni değil, hepi topu az buçuk xml serialization. görselleştirme aracım ek$iVista da güncellemeden payını alacak. ilk sürümünde directed graph çizebilmek için netron kullanmıştım. kendisi artık bir ticari ürün ve bu dönüşümü geçirmeden önceki halinden de pek memnun değildim. iyi bir paket ararken microsoft research'tan c#ung'u buldum. "chung" diye okunuyormuş kendisi. paketin içinden birkaç dll, dokümantasyon ve excel add-in'i geliyor. bu add-in çok hoşuma gitti; iki sütuna directed graph'ın edge'lerinin başlangıç ve bitiş vertex'lerini sırayla yazıyorsun ve araç çubuğundan c#ung'u çağırıyorsun, grafik hemen karşında. layout algoritması olarak da dairesel, fruchterman-reingold, sugiyama ve grid algoritmaları hazır geliyor. vertex'leri extend edip biraz interaktivite ekledik mi işimi görür bu c#ung, her ne kadar java dünyasındaki karşılığı (ve öncülü) jung kadar gelişkin olmasa da.

önceki gibi ek$iVista da splash screen görüntülenirken ön yükleyici çalışacak ve veritabanındaki kaynak başlıkları (başka başlıklara link içeren başlıklar) bir veri yapısına alınacak. c#ung'a vereceğim kaynak-hedef çiftlerini veritabanından almak için de bir sql fonksiyonu yazdım; hangi başlıktan kaç tık mesafeye kadar gidileceği girdi olarak alınıp kaynak-hedef çiftleri döndürüyor. önceki versiyonumuzda yoktu bu ve bu yüzden grafik çizimi sırasında sorgular dallanıp budaklandığı için makina fena kastırıyordu.

bu işin tek geliştiricisi olsam da arada kodu fena halde boklayabildiğim için bir versiyon kontrol sistemine geçmek gerekti. ben subversion'da karar kıldım, du bakali nolcek...

benden ilgi bekleyen işlerin bir listesini yapayım dedim, listenin altında kaldım. ahanda:
  • ek$iVista revisited (yukarıdaki kalabalık)
  • byblos kütüphane/exlibris mevzuları
  • muha! kişisel muhasebe dalgametresi
  • plone cms incelemeleri
  • -- confidential --
  • altivi analiz/listener/falan
  • quant test/anket cihazı
  • nonlinear forum
  • versaTile erp/simulasyon/jack-of-all-trades
  • inanclisesi.net facelift (plone incelemesi ile bağlantılı)
  • r/c autonomous cihazlar araştırma, diy işleri
  • mezunlar derneği çalışmaları
  • girişimciler kulübü
çok çalışmam lazım...

buradan sonrası da yukarıdaki listeyle alakalı. hım hım hım hım, eveeet:
  • ek$i olayına yukarıda detaylıca değindik, tekrara lüzum yok.
  • byblos'ta veri yapısını kurmakla uğraşıyorum.
  • muha! için biraz muhasebe öğrenmem gerekiyor. fon/faiz gelirlerini, kredileri, kredi kartlarını, senetleri falan nasıl muhasebeye yansıtacağım hakkında hiçbir fikrim yok.
  • plone gayet temiz bir cms. incelemye fırsatım pek olmasa da birçok da eklentisi var. genişletilebilir her şey iyidir.
  • altivi analiz için halen saçmasapan bir dikdörtgenler prizmasını nasıl çizeceğimi bulmaya çalışıyorum. saçmasapan, çünkü daha önce öngörmediğim bir durum var. bir otomobil ihalesinin inceliyoruz diyelim; tavan fiyat onbinlerce ytl, yani en kötü durumda fiyat ekseni üzerinde onbinlerce görülebilir boyutta olması gereken birim olacak. bu durum diğer iki ekseni incelenemez kılabilir, böyle olunca da grafiği bu formatta hazırlamanın bir anlamı kalmıyor. galiba asansörlerden çok buna kafa yormak lazım:P altivi listener ise ek$iAPI işini bekliyor, çünkü altivi'den veri toplama işini ek$iAPI'ı kullanarak yapıyorum.
  • quant ve versaTile konusunda herhangi bir eylem planım yok, şimdilik uykudalar.
  • nonlinear forum daha pişmedi.
  • inanclisesi.network olayını okul yönetimiyle paylaşmayı düşünüyorum. konu ile ilgili birkaç ilginç olduğunu düşündüğüm fikrim var 3d-mıridi falan, ama önümdeki dikdörtgenler prizmasını halledeyim ben önce sanki.
  • mezunlar derneği "türk eğitim vakfı inanç türkeş özel lisesi mezunları derneği" olarak değil de "gebze inanç türkeş özel lisesi mezunları derneği" olarak kuruluyor. belgeler bugün il dernekler müdürlüğü'ne iletildi, kısa sürede çalışmaya başlayabilirmişiz gibi geliyor bana yoksa şüphem mi var?
  • girişimciler kulübü işi caner lojmanına ve ortamına iyice yerleşip alışınca başlayacak umarım. aceleye de getirmiyoruz, memleketi kurtarmak uzun, zorlu bir süreç ve ülkemin kahvehanelerinde başlıyor. efkarın aşama kaydetmeye engel olan bir etken olduğu da buralarda keşfedildi, ki o yüzden kimse kahvehane safhasından ileri gidemiyor. biz bir ihtimal gidebileceğimizi düşünüyoruz.
  • r/c otonom araçlar işi ise darpa grand challenge ile ilgili bir video izleyince takıldı kafama. tabi bir hummer alıp milyon dolar harcamayacağım, ama buna göre daha mikro düzeyde kalan şeyler yapabilirim. ne bileyim, orta büyüklükte uzaktan kumandalı araçlara microcontroller yerleştirip delicesine hack etmek istiyor make etkisindeki bünye. of ya, of! patates bazukası daha kolaydı sanki...

dün sezai bey'i ziyaret ettik. azdık, ama olsun. nur içinde yat...

Thursday, July 5, 2007

mcinfaaoals - eklemeler

madde madde kasayım olan bitenleri:

  • geçen hafta servet'le buluştum, ki onunla görüş(e)meyeli -hiç abartmayayım- üç yıl oluyor. uzakta değilmiş en azından; ben iş kuleleri'ndeyim, o da yapı kredi plaza'daki kpmg istanbul ofisinde. yürüme mesafesi yani. artık daha sık görüşecek olmamız güzel birşey, ama bu konuda servet ne düşünür bilemeyeceğim :P arada, daha doğrusu oraya buraya teftişe çıkmadığı zamanlar da akbank'ta müfettişlik parçalayan erkal da bize katılır herhalde. sözün özü, levent havalisinde ufak bir inanç kliği oluştu, bu civara uğrayanları bekleriz.
  • altivi teklif detayları sömürgeni bir uygulama hazırlıyordum makina çökmeden önce, onu tamamladım. acayip kalabalık bir veritabanı oluştu; şimdilik yaklaşık 3800 ihale, 650-700 bin kadar da teklif var incelenecek. ilk izlenimlerim biraz şaşırtıcı, hafiften işkillendirici, sonrasında da fena halde gıcık edici.
    • şaşırtıcı olan kısmı, eğer altivi sattığı ürünleri "ürün fiyatı" diye duyurduğu fiyattan temin ediyorsa, şu anki teklif sayısı ve teklif bedelleri göz önünde bulundurulduğunda yaklaşık 2 milyon ytl kadar zararda görünüyor. tabii ki, bu gayet naif bir varsayım; yani piyasa fiyatlarından biraz daha düşük fiyatlarla satınalım yapıyor olmaları beklenir. altivi'den mi yoksa başka bir yerdenmi okudum çok iyi hatırlamıyorum, ama stok tutmadıkları şeklinde bir bilgi/duyum var ki, toptan alım yapmadan nasıl indirim alınabilindiğini çözemedim. indirim demişken, perakende fiyatı üzerinden yaklaşık %25 indirim almış olmaları gerekiyor, o da başa baş noktasına gelinmesi için. bu durumda ya stok tutuyorlar, ya da distribütörlerden acayip kelepir mal kapatıyorlar.
    • işkillendirici olan, tekliflerin önemli bir kısmının son dakika içinde yapılmış olması, hatta birçok durumda tekliflerin neredeyse kazanan rakamı ya merkez alacak şekilde, ya da ucu ucuna içerecek aralıklar içinde verilmesi. eğer verdiğiniz teklif grubundan daha sonra teklif verilmeyeceğinden eminseniz o zamana kadar verilen teklif sayısından 1 fazla teklifi en yüksek fiyattan aşağı doğru verirseniz kazanmanız garanti, ama bu da çoğu zaman karlı olmuyor. esas işkillendirici olan ise tekliflerin veriliş şekli değil, bunların altivi ekibi tarafından verilen dummy, "keriz silkeleme" amaçlı teklifler olma ihtimali. bu bakımdan biraz şeffaflığa ihtiyacı var altivi'nin.
    • gıcık edici kısım kendini sitenin "yasal uyarı" kısmında gösteriyor. deniyor ki:
      AL SATIŞ bu internet sitesinin genel görünüm ve dizaynı ile internet sitesindeki tüm bilgi, resim, AL TİVİ markası ve diğer markalar, www.altivi.com ve www.altivi.com.tr alan adları, logo, ikon, demonstratif, yazılı, elektronik, grafik veya makinede okunabilir şekilde sunulan teknik veriler, bilgisayar yazılımları, uygulanan satış sistemi, iş metodu ve iş modeli de dahil tüm materyallerin (“Materyaller”) ve bunlara ilişkin fikri ve sınai mülkiyet haklarının sahibi veya lisans sahibidir ve yasal koruma altındadır.
      bold italic kısım beni değil gıcık, resmen ifrit etti. takası icat eden adam patentini almayarak aptallık mı etmiş? bu işi ilk sen mi yapmışsın, mucidi sen misin? hepsini geçtim, bir satış yöntemi herhangi bir şekilde fikri koruma altına alınabilir mi? bana göre tüm bu soruların cevabı hayır. öncelikle bu iş altivi'cilerin gri hücrelerinin mahsulü değil, kendilerinden önce açılmış örnekler var (mesela http://www.limbo.com/), satış yöntemi de çok taze değil, o halde pazarı kapatmak için böyle bir cinliğe başvuruyoruz. ayıp değildir de nedir yani şimdi bu? incelemelerimiz sürecek; e, soruşturmacı gazeteciliğin tadını aldık bir kere, durmak olur mu :P
  • her bir procem beklemede, ama artık suçu zamansızlığa değil de maymun iştahlılığa bağlıyorum artık. bir onunla uğraşayım, bir de şuna bakayım derken hiçbiriyle hakkıyla ilgilenemiyorum. işleri bir öncelik sırasına koyup teker teker ele almak en doğrusu...
  • gödel, escher, bach: an eternal golden braid... hastası olduğum, her elime alıp sayfalarını karıştırdığımda mutlaka bir şekilde beni şaşırtan, afallatan, aynı anda hem daha zeki ve daha aptal, eksik hissettiren bir kitap. koç'tayken orijinalinden, ucundan kenarından nasiplenebilmiştim. sonra kabalcı'dan türkçe çevirisinin çıktığını öğrendim ve bir tane edindim. çevirisi fena sayılmaz ve kesinlikle öneririm, özellikle temel bilimciler ve mühendislere. tabii, imkanınız varsa orijinalinden, doğrudan douglas hofstadter'in elinden çıkma metni takip etmeniz daha uygun, daha güzel olur.
  • işte ise yepyeni bir macera: "iş bankası outsourcing öğreniyor". accenture ile çalışıyor, onların manila'daki ekibine işleri usulünce yapabilsinler diye deli gibi doküman hazırlıyoruz. hatta geçen gün 2 saat telekonferansla elemanlara belli tip bir servisin nasıl yazılacağıyla ilgili sunum yaptım. hiç yapmadığım şey... neyse ki iyi gitti. işin garibi, yazılım departmanında kod yazmayı özledim, hazır -ve ne yazık ki köhnemiş- codebase üzerinde at koşturmak yerine yeni bir şeyler yapmayı özledim.
biriktirip biriktirip patlıyoruz işte böyle efendim.

Friday, June 23, 2006

comp 491: report | conclusion

Conclusion

This project, while it only consisted of ek$iAPI, had started as an exercise in C#, not thinking that a senior design project could be based on it. The source of inspiration for this graph visualization application was the Skitter(*) project which also featured a circular graph, but laid out in a completely different fashion.

After making the decision of doing this project, a tremendous amount of effort was expended to complete it. However, even more could have been expended for a total fulfillment. Quoting from the Preliminary Report:

Scope: A detailed inspection of Ekşi Sözlük data in the form of a digraph as a way of representation, with some simple algorithms employed for coming up with the digraph. Extensions, such as marking the titles one specific suser has written, finding cycles of association or creating timelines (or a histogram) of activity for a specific title can also be implemented.”

“The latter and final step is to design and implement the graphing tool which will work on the extracted data. This tool will make use of some simple algorithms or checks. Some are:
  • Checking the number of entries under a destination title before assigning a connection between two nodes depicting titles. This will be necessary, as links sometimes are used for other purposes by susers, such as emphasizing a part of the entry. Also, some links point to non-existent titles which should be eliminated.
  • Possibly, a node distribution algorithm, so that no node of the graph overlaps with another to allow clarity of presentation.”

The prime objective of the project can be said to be accomplished, as a digraph is generated by ek$iVista. There is a very, very simple algorithm to come up with the digraph; no need was seen for checking the number of entries under a destination title as that quantity carries no importance. References pointing to single entries and clever references were left out, because the connections sought have to be between titles and clever references are generally used to make remarks about a fact and carry no little referential value. The envisaged extensions that were left out in the first version of ek$iVista are implemented in the second, such as the activity histogram or a list of common titles of two arbitrary susers.

Of course, there is plenty of room for improvement. Edges or vertices could be colored according to a measure, such as the number of links from the edge, or susers in the ”yazarlar” tab could be assigned different icons according to their generations

As I stated above, this project started as a small exercise in C# language and expanded into a much bulkier one, helping me master very crucial constructs; accessing databases, acquiring data from the Internet, working with basic graphics, using proprietary packages and many other skills.

Hoping that somebody comes up with a programming language exercise that also improves one’s time management skills…


K. Egemen Şentin
27.01.2006, updated 11.02.2006



(*)Website: http://www.caida.org/analysis/topology/as_core_network/

comp 491: report | ek$iEdgeDump

ek$iEdgeDump

The Problem: Extracting links (connections) from the existing “heap” of entries.

Design: For ek$iVista to be able to function, to be able to produce a digraph, it needs a list of directed edges, and ek$iEdgeDump was produced for this purpose. For the sake of simplicity, like ek$iDump, it is also designed as a console application. During the development, two versions of ek$iEdgeDump were produced. The first version gets the full list of titles in the database (a table named Titles exists in the database) and scans them one by one. In this scan, the entries under the title being scanned are inspected and any link that points to a title (those that point to single entries are omitted) is parsed out of the entry text. Then, another database query checks whether the title pointed by the link exists in the title list. If it exists, the pair consisting of the IDs of the source title and the destination title (the records in the Titles table have a title ID and title name) are written to a table named EdgeData, only to be used by ek$iVista in drawing the digraph. This approach proved to be too slow, because for every link found in an entry, a verification query has to be made. The scan rate of this version of ek$iEdgeDump was less than 1,000 titles/day. Given the fact that the database contained more than 700,000 titles, the job would be completed in nearly two years. Clearly, another approach had to be adopted .

In the second version of ek$iEdgeDump, the focus is back on the entries instead of the titles. As one can recall from the description of the Entry class in ek$iAPI, one of the details of acquired from Ekşi Sözlük when an entry is extracted is the title the entry is placed under. Thus, we can produce a different table that looks like the EdgeData table described above that keeps information of the source and the destination vertices of the directed edge. The table, in the new approach, is produced by scanning the entries in the database (they reside in a table named Entries), parsing out the links that point to titles and writing the pair consisting from the name of the source title and the destination title to a table named EksiEdgeData without checking whether the destination title exists in the Titles table. This verification effort was the factor that slowed the first version down, and it can be handled without querying the database by ek$iVista (the details of how this is done are given in the section discussing ek$iVista). As the title data in the Entries table is stored in string format (not as integers; foreign keys related to title ID column in Titles table), the size of the EksiEdgeData table is significantly larger than that of EdgeData. ek$iVista uses the data from EksiEdgeData table, generated by the last version of ek$iEdgeDump.




Fig. 5: ek$iEdgeDump versions 1 and 2 in action

comp 491: report | ek$iDump

ek$iDump

The Problem:
Coming up with a portable application to acquire Ekşi Sözlük entries and store them in a database.

Design: Although the final product of the project will be a graph depicting connections between Ekşi Sözlük titles, the connections arise from the content of the titles, which are, obviously, the entries. That is one of the reasons why ek$iDump is a tool for getting the entries rather than the titles. Another and maybe the prime reason for focusing on entries is that entries have unique integer IDs that allow them to be acquired one by one in a for-loop or a while-loop.

ek$iDump is, due to this nature of Ekşi Sözlük entries, at the level of complexity of a “Hello World” program. The program, designed as a console application, gets the starting ID and the terminal ID as its input, which are integer values. In a while-loop, beginning from the starting ID, if the entry with the given ID exists, it gets it from Ekşi Sözlük by calling Entry.GetFromEksi(ID) and writes the details of the Entry acquired to the database. The database of choice is a Microsoft Access file, because one does not have to set up a server for using it; even if you do not have Microsoft Access installed, one can obtain and install a package named Office 2003 Redistributable Primary Interop Assemblies and get on with using the database. Also, the data accumulated in the database is easily exportable to Microsoft SQL Server, which the other two sections of the project, ek$iEdgeDump and ek$iVista use. One always has to make the quantum leap from Microsoft Access to another pro-level RDBMS at some level, as Microsoft Access imposes a size limit of 2 GB on a database file. Note that although it is by no means final, the size of the database file generated in the course of the project exceeds 5 GB. An image showing ek$iDump in action is given in the figure below:


Fig. 4: ek$iDump in action, dumping entries

comp 491: report | ek$iAPI

ek$iAPI

The Problem: Retrieving information from Ekşi Sözlük and organizing it programmatically.

Design: In the very beginning, the plans were basically getting started with ek$iDump and incorporating data acquisition functions into some method inside the main class. That would have been very easy to start with, however, if I wanted to use the portions of code that access Ekşi Sözlük in another application, all had to be re-written. To avoid such circumstances, a different style of programming had to be adopted. After some contemplation, I decided to program this part of the project as a class library. Class libraries are collections of classes and methods inside classes, and when compiled in Visual Studio.NET, are built into Windows DLLs. Thus, I would be able to reuse the code in any application.
Although it carries no formal significance, ek$iAPI did spring from the drawing board:


Fig. 3: Primordial sketches


In the initial sketch, there number of classes envisaged was six; entry class for storing entry data, a message class to exploit the messaging facility of Ekşi Sözlük, a today class for storing the titles written to in a specific day, a random50 class for returning 50 titles chosen at random by using the search facility, a suser class to store any particular information obtainable about an Ekşi Sözlük user, and a title class for storing entries posted under a title. In the final release, however, some of these envisaged classes were dropped. Message class was discarded as it would provide no functionality for the time being; classes today and random50 were discarded as they were collections of title objects and thus were redundant. As a result, the only classes implemented which were also on the “primordial sketch” are Entry, Title and Suser classes.

This class library is utilized by ek$iDump, Ekşi Sözlük entry acquisition tool, which is described in the next section.

comp 491: report | project divisions

Project Divisions

Ekşi Sözlük graph visualization project consists of four distinct applications:
  • ek$iAPI: Class library (basically, a Windows DLL) for acquiring, manipulating and organizing Ekşi Sözlük data, coded in C#
  • ek$iDump: A console application coded in C# which exploits ek$iAPI to retrieve entries from Ekşi Sözlük and “dump” them into an MS Access Database
  • ek$iEdgeDump: A console application coded in C# which processes the entries acquired by ek$iDump and finds links between titles
  • ek$iVista: The final application that outputs an image file depicting the connections between titles in Ekşi Sözlük using the data generated by ek$iEdgeDump.
Design and implementation details are given in the following sections.

Thursday, June 22, 2006

comp 491: preliminary report

hani proce raporu diyorduk ya, işte ta kendisi. ancaaaaak, önce neyin üzerine rapor yazıyoruz bilelim, di'mi?

i kept on telling about some project report, and there it is. but first of all, we have to know what this report is written about, innit?


12.11.2005

Topic:
Ekşi Sözlük Graph Visualization Tool

Motivation: Ekşi Sözlük (http://sozluk.sourtimes.org/) is a popular Turkish web site, up and running since February 15th, 1999. Having about 10,000 active contributors (susers – Sözlük users in Ekşi Sözlük jargon), this web site is basically a hypertext dictionary comprising of the entries of its collaborators. In Ekşi Sözlük, one can find explanations and definitions of almost any concept one can think of. In Ekşi Sözlük’s jargon, a concept for which information can be found is called a “title” (literal translation of “başlık” from Turkish). Each individual definition, explanation, or information of any kind is called an “entry”. There may be any number of entries posted under a title. What makes Sözlük different from any other plain text based dictionary is that it contains hyper-textual references to other titles. The data to be used in this project is obtained by crawling through the entries of Ekşi Sözlük.

Scope: A detailed inspection of Ekşi Sözlük data in the form of a digraph as a way of representation, with some simple algorithms employed for coming up with the digraph. Extensions, such as marking the titles one specific suser has written, finding cycles of association or creating timelines (or a histogram) of activity for a specific title can also be implemented.

Method: As an initial step, a crawler for extracting Ekşi Sözlük data, named ek$iDump was written in C#, which is a simple, single-threaded application which accesses Sözlük entries one by one by their numerical ID and dumps the necessary details to a non-relational Microsoft Access database. Currently all entries until the ID #3300000 have been crawled. Due to the high number of deleted entries by moderation, the choice of the suser or voiding of the suser account, the total number of entries stored locally stand close to 2,000,000. As of December 12th, 2005, there are more than 5,000,000 entries posted under about 1,100,000 titles and the ID of the most recent entry is #8685844. This may give a measure of the density of Sözlük data (detailed statistics can be found at http://sozluk.sourtimes.org/stats.asp). Due to time limitations, a cutoff point will be selected (ID #4000000 or #5000000 is considered). The latter and final step is to design and implement the graphing tool which will work on the extracted data. This tool will make use of some simple algorithms or checks. Some are:
  • Checking the number of entries under a destination title before assigning a connection between two nodes depicting titles. This will be necessary, as links sometimes are used for other purposes by susers, such as emphasizing a part of the entry. Also, some links point to non-existent titles which should be eliminated.
  • Possibly, a node distribution algorithm, so that no node of the graph overlaps with another to allow clarity of presentation.
Expected Results: A report of the senior design project with extended demonstrations of the final product, the graphing tool which is expected to generate a “forest” of Ekşi Sözlük data. As mentioned above, the data extraction tool (crawler) is complete with a collection of classes to be able to acquire and arrange Ekşi Sözlük data, namely the Ek$iAPI; although can still be improved speedwise. The graphing tool is currently in the drawing-board phase.